Papers with systematic evaluation of

5 papers
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents (2026.acl-long)

Copied to clipboard

Challenge: Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, but systematic evaluation of GUI–shortcut hybrid agents remains underexplored.
Approach: They propose a benchmark that evaluates GUI-shortcut hybrid agents with a specific focus on the mobile domain.
Outcome: MAS-Bench evaluates agent's ability to generate shortcuts by discovering and creating reusable, low-cost workflows.
InferES : A Natural Language Inference Corpus for Spanish Featuring Negation-Based Contrastive and Adversarial Examples (2022.coling-1)

Copied to clipboard

Challenge: InferES is an original corpus for Natural Language Inference (NLI) in European Spanish .
Approach: They propose to implement and analyze a corpus-creating strategy utilizing expert linguists and crowd workers to provide high-quality data and facilitate the systematic evaluation of automated systems.
Outcome: The proposed model obtains 72.8% accuracy and performs moderately well on negation-based adversarial examples.
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Retrieval-Augmented Generation allows to enhance Large Language Models with external knowledge.
Approach: They propose a library that allows to benchmark and standardize RAG experiments.
Outcome: The proposed library is an end-to-end library for reproducible research standardizing RAG experiments.
Hallucination Detection in Structured Query Generation via LLM Self-Debating (2025.findings-emnlp)

Copied to clipboard

Challenge: Hallucination remains a key challenge in applying large language models to structured query generation . we propose the Self-Debating framework to enhance detection performance .
Approach: They propose a framework that prompts an LLM to generate contrastive explanations from opposing perspectives . they also propose 'self-debating' framework to enhance detection performance .
Outcome: The proposed framework outperforms LLM-as-a-Judge baselines in hallucination detection . the framework generates contrastive explanations from opposing perspectives .
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls.
Approach: They introduce a diagnostic benchmark and conduct a systematic evaluation of multilingual tool calling across Chinese, Hindi, and the low-resource language Igbo.
Outcome: The proposed benchmarks show that multilingual tool calling fails despite correct intent understanding and tool selection.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations